Papers with data cleaning

11 papers
A Confidence-based Acquisition Model for Self-supervised Active Learning and Label Correction (2025.tacl-1)

Copied to clipboard

Challenge: Existing approaches to training deep neural networks require large amounts of meticulously annotated data.
Approach: They propose a pool-based active learning framework that requires expert annotators to label only a fraction of a sequence and facilitates self-supervision for the remainder of the sequence.
Outcome: The proposed model outperforms baselines on dialogue belief tracking tasks.
Sailor: Open Language Models for South-East Asia (2024.emnlp-demo)

Copied to clipboard

Challenge: Large language models (LLMs) rely on English data for training, but are often not comparable across other languages.
Approach: They propose to develop a family of open language models for SEA languages . they use BPE dropout, aggressive data cleaning and deduplication to improve model robustness .
Outcome: The proposed models perform well across four benchmarks, including commonsense reasoning, question answering, reading comprehension and examination.
Advancing E-commerce Merchants Telemarketing with Synthetic Data-Driven LLMs (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are proving broadly applicable across diverse industries, including e-commerce.
Approach: They propose a hybrid data synthesis framework that unifies the input schema with profile and strategy designed by top sales and extracts them via a Multi-task paradigm.
Outcome: The proposed model reaches the performance level of the top 25% of human sales in terms of the final marketing results.
A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solution (2025.emnlp-main)

Copied to clipboard

Challenge: Low-resource language understanding is challenging for large language models (LLMs).
Approach: They propose a CompRehensive lIterary Chinese readIng comprehenSion procedure with a large dataset for CRISIS.
Outcome: The proposed procedure has the largest dataset and substantiates the effectiveness of the proposed procedure with a 7 percent hike in accuracy compared with the baseline.
Ask Language Model to Clean Your Noisy Translation Data (2023.findings-emnlp)

Copied to clipboard

Challenge: Neural machine translation models exhibit a noticeable decline in translation quality when exposed to noisy input.
Approach: They use a dataset to evaluate the robustness of NMT models against noisy inputs.
Outcome: The proposed dataset cleaners the noise from the target sentences while preserving the semantic integrity of the original sentences.
Tilde MT Platform for Developing Client Specific MT Solutions (L18-1)

Copied to clipboard

Challenge: a growing demand for translations and multilingual content is surpassing the supply of professional translation services.
Approach: They present a custom machine translation platform called Tilde MT that provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality.
Outcome: The proposed platform provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality, and wide integration capabilities.
MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation (2025.coling-main)

Copied to clipboard

Challenge: Existing datasets suffer from limitations in scale, diversity, and quality, hindering the development and evaluation of IT models.
Approach: They propose a large-scale parallel corpus of multilingual image translation with over 10M image-text pairs derived from real-world data.
Outcome: The proposed model performs better in tackling challenging and complex image translation tasks in the real world.
Make Every Example Count: On the Stability and Utility of Self-Influence for Learning from Noisy NLP Datasets (2023.emnlp-main)

Copied to clipboard

Challenge: Increasingly larger datasets have become a standard ingredient to advancing the state-of-the-art in NLP, however, data quality might have already become the bottleneck to unlock further gains.
Approach: They propose a general method for improving model performance in the presence of noisy training data based on self-influence and bandit curriculum learning.
Outcome: The proposed method improves model performance in machine translation, question answering and text classification, building up on approaches to self-influence calculation and automated curriculum learning.
Can Large Language Models Fix Data Annotation Errors? An Empirical Study Using Debatepedia for Query-Focused Text Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Debatepedia dataset limited by noise and most queries do not have relevance to document .
Approach: They harness the language generation capabilities of two LLMs to regenerate queries in a Debatepedia dataset.
Outcome: The proposed model can regenerate queries from the Debatepedia dataset.
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have advanced the automation of data science workflows, yet it remains unclear whether they can critically leverage external domain knowledge as human data scientists do in practice.
Approach: They propose a benchmark to evaluate how large language models handle external domain knowledge in tabular prediction tasks.
Outcome: The proposed model evaluates whether it can critically leverage external domain knowledge as human data scientists do in practice.
Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Language models such as GPT and Llama have shown remarkable ability on diverse natural language tasks, yet their performance on complex table tasks is suboptimal.
Approach: They propose a generator-validator paradigm to iteratively generate-then-validate training data from language models to fine-tune stronger Table-Specialist models that can specialize in a given task, without using manually-labeled data.
Outcome: The proposed model outperforms vanilla language models on diverse table tasks and can match or surpass GPT-4 level quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations